Skip to content

perf(io): parallel order-preserving BGZF compression for BAM output - #276

Open
BenjaminDEMAILLE wants to merge 2 commits into
mainfrom
perf/parallel-bgzf-writer
Open

BenjaminDEMAILLE wants to merge 2 commits into
mainfrom
perf/parallel-bgzf-writer

Conversation

@BenjaminDEMAILLE

Copy link
Copy Markdown
Contributor

Refs #223 (option 1: parallel compression; record encoding still single-threaded).

Wall time, 2M SE 100 bp reads, synthetic 20 Mb genome, BAM Unsorted, 16-core Mac (median of 3):

threads no output BAM before BAM after
1 12.35s 12.88s 12.94s
4 5.17s 5.35s 5.42s
8 2.87s 3.08s 2.97s
12 2.30s 3.40s 2.54s
16 2.32s 4.18s 2.83s

At 16 threads with --outBAMcompression 6: 5.00s → 2.92s.

Follow-up (not here): move BAM record encoding into the parallel align stage (conflicts with #222/#261), and bam_dedup.rs's own writer.

Tests: fmt, clippy 0 warnings, cargo test 636 passed.

🤖 Generated with Claude Code

New BgzfWriter: the caller fills 64 KiB blocks, --runThreadN dedicated
threads compress them with libdeflate, and one writer thread emits them in
submission order. Output is byte-identical to noodles_bgzf at every thread
count. All four BAM writers use it. finish() now writes the EOF marker and
propagates errors instead of relying on drop.

Refs #223

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Merge of origin/main into #276: only src/io/bam.rs conflicted (main's CoordinateSorter external sort vs #276's BgzfWriter).
Resolved by taking main's bam.rs and re-applying #276: make_bgzf_writer(inner, level, threads) -> BgzfWriter, bgzf_threads(params), threads threaded through BamWriter/BamStdoutWriter/CoordinateSorter (spill runs and reduce passes also compress in parallel), finish via get_mut().finish(). Sort semantics, tie order and memory bound are main's; public API shape of #276 unchanged (BamWriter::with_header gains threads).
Verified (yeast 200k pairs, 8 threads): samtools view bodies identical to a main build for BAM Unsorted, SortedByCoordinate, and SortedByCoordinate --limitBAMsortRAM 1000000 (spills). cargo fmt, clippy -D warnings, test --release pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
BenjaminDEMAILLE added a commit that referenced this pull request Oct 9, 2026
…l-record-encoding

Merge of origin/perf/parallel-bgzf-writer (#276: main's external CoordinateSorter + parallel BGZF) into #278 (worker-side SAM/BAM encoding).
Conflicts: src/io/bam.rs (7), src/io/sam.rs (1), src/lib.rs (2). sam.rs: kept #276's Default-based BufferedSamRecords::new with #278's encoded fields.
bam.rs: CoordinateSorter now buffers BAM-encoded bytes + (ref,pos,off,len) index (replaces #278's SortBuffer and the RecordBuf buffer), spills sorted runs as headerless BGZF of raw records, k-way merges raw records (ties: input order via stable sort + run index). Budget counts real bytes + index entries (estimated_record_bytes removed; its test replaced). SortedBam(Stdout)Writer keep encoder()/write_encoded().
lib.rs: main now builds WithinBAM records on the worker into buf.records, so #278's writer-side supplementary writes and PreEncoder.within_bam were dropped.
Verified: fmt, clippy -D warnings, cargo test --release pass; yeast sidx 200k pairs, 8 threads: SAM / BAM Unsorted / Sorted / Sorted+limitBAMsortRAM 1000000 samtools-view bodies md5-identical to a build of origin/perf/parallel-bgzf-writer (with and without RUSTAR_PRE_ENCODE=1).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant